Papers by Mohammad Ruhul Amin

6 papers
A Study on Using Semantic Word Associations to Predict the Success of a Novel (2021.starsem-1)

Copied to clipboard

Challenge: Existing methods for book success prediction are not effective.
Approach: They propose to represent a book as a spectrum of concepts based on the association score between its content embedding and a global embeddment for a set of semantically linked word clusters.
Outcome: The proposed method outperforms the previous methods for book success prediction.
BD-SHS: A Benchmark Dataset for Learning to Detect Online Bangla Hate Speech in Different Social Contexts (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS) due to the massive amount of user-generated content, modern machine learning techniques are feasible and cost-effective to tackle this problem.
Approach: They propose to use a large manually labeled Bangla HS dataset to train generalizable models.
Outcome: The proposed dataset includes more than 50,200 offensive comments crawled from online social networking sites and is at least 60% larger than existing Bangla HS datasets.
BanSuite: A Unified Toolkit and Software Platform for Low-Resource NLP in Bangla (2026.eacl-demo)

Copied to clipboard

Challenge: Existing efforts to improve Bangla's NLP performance have focused on isolated tasks such as Part-of-Speech tagging and Named Entity Recognition (NER) but comprehensive, integrated systems for core NLP tasks such Shallow Parsing and Dependency Parser are largely absent.
Approach: They propose to integrate a large-scale, manually annotated Bangla Treebank with high-quality pretrained models for POS tagging, NER, shallow parsing, and dependency parse.
Outcome: The proposed system achieves strong in-domain baseline performance while maintaining high efficiency in resource usage.
AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning for Small and Large Language Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models excel at complex reasoning tasks, yet their performance hinges on the quality of their prompts and pipeline structures.
Approach: They propose a framework that fully automates large language models' pipeline construction using reinforcement learning.
Outcome: Experimental results show that autoDSPy outperforms DSPy benchmarks in accuracy gains and time.
SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts (2021.findings-emnlp)

Copied to clipboard

Challenge: Bangla is the sixth most spoken language worldwide and the second Indo-Aryan language after Hindi.
Approach: They propose an annotated sentiment analysis dataset made of informally written Bangla texts.
Outcome: The proposed dataset is compared with neural networks and pretrained models . it shows that hand-crafted lexical features provide superior performance than neural networks .
BanNERD: A Benchmark Dataset and Context-Driven Approach for Bangla Named Entity Recognition (2025.findings-naacl)

Copied to clipboard

Challenge: In a cross-dataset evaluation, models trained on BanNERD consistently outperformed those trained on four existing Bangla NER datasets.
Approach: They propose to use Bangla as a language to create the most extensive human-annotated and validated Bangla NLP dataset.
Outcome: The proposed method outperforms existing methods on Bangla NER datasets and performs competitively on English datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations